Gradient matching attack method of Vision Transform model for noise disturbance
By constructing an attention-aware gradient diffusion denoising network and ViT self-attention mechanism, the gradient leakage problem of existing methods under noise perturbations in the Transformer model is solved, and high-fidelity gradient recovery and privacy protection are achieved.
Patent Information
- Application Number
- CN202510784406.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
AI Technical Summary
Existing gradient leakage attack methods cannot effectively target Transformer models, especially under noise perturbations, and cannot effectively restore high-fidelity original semantic information, resulting in insufficient privacy protection.
By constructing an attention-aware gradient diffusion denoising network, the diffusion model is used to gradually remove the noise components, and combined with the ViT self-attention mechanism, a gradient matching attack is performed to restore the original semantic information.
A high-fidelity gradient leakage attack is implemented on the noise-perturbed Vision Transformer model, enhancing privacy protection in federated learning scenarios.
Smart Images

Figure CN120688088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the issue of privacy protection in federated learning scenarios, and more specifically to a gradient leakage attack method for a noise-perturbed Vision Transformer model. Background Art
[0002] Gradient updates are a core element of federated learning and hold significant application value in privacy-sensitive fields such as medical image analysis, financial risk control modeling, and smart city perception. In the field of data security, gradient perturbation through the addition of noise has become a mainstream technique for protecting user privacy. Its theoretical foundation, built on the framework of differential privacy, effectively defends against traditional gradient matching attacks. However, with the evolution of adversarial attack techniques, attackers, by constructing deep generative networks and optimization algorithms, are able to extract sensitive data features with semantic information from noisy gradients, posing a serious threat to federated learning systems that rely on gradient sharing. Existing defenses often employ static noise addition methods, such as Gaussian noise. While these methods can enhance data anonymity, they struggle to balance privacy protection with model convergence efficiency. Especially when processing high-dimensional, sparse data, noise injection can significantly degrade model accuracy. In medical image analysis, patient pathology slides contain high-resolution biometric features, allowing attackers to reconstruct recognizable lesion images through multiple rounds of correlation analysis of gradient sequences. In the field of financial credit assessment, user consumption behavior data exhibits strong temporal correlations, and noisy gradients can be used to reverse-infer transaction patterns using hidden Markov models. Even more challenging, IoT edge devices are limited by computing power and cannot implement dynamic adaptive noising strategies, allowing attackers to extract gradient feature correlations through device fingerprints. Therefore, it is urgent to study attack methods against noisy gradients from an adversarial perspective to enhance privacy protection in federated learning scenarios.
[0003] Existing gradient leakage attacks can be roughly divided into two directions: gradient analysis attacks and gradient matching attacks. The paper "Deep Leakage from Gradients (DLG)" was the first to achieve the recovery of original data from the gradient information generated by model training. However, when faced with gradients protected by noise, DLG's ability to reconstruct and recover data is significantly weakened. The paper "Mjolnir: Breaking the Shield of Perturbation Protected Gradients via AdaptiveDiffusion" attempts to use the diffusion model to perform gradient leakage attacks. Mjolnir first denoises the noisy gradients based on the diffusion model, and then reconstructs the denoised gradients using DLG. However, the method proposed in the paper Mjolnir is only applicable to simple convolutional neural networks, such as LeNet. Currently, the Transformer model is becoming more and more widely used. Existing gradient leakage attack methods cannot be directly transferred to the Transformer due to differences in model architecture. There is a lack of gradient leakage attack methods specifically targeting Transformer models with noise perturbations.
[0004] Therefore, the present invention proposes a gradient denoising attack method based on a diffusion model for the noise-perturbed Vision Transformer (ViT) model. In response to the feature distortion problem caused by traditional Gaussian noise, the present invention designs an attention-aware gradient diffusion denoising network, which reconstructs high-fidelity original semantic information from the noisy gradient by implicitly modeling the temporal correlation of the gradient sequence and the feature correlation of the attention head. Specifically, by constructing a conditional diffusion model, the noise components injected into the gradient are gradually stripped off during the diffusion process, and based on the interactive characteristics of the class token and the image patch in the ViT self-attention mechanism, a gradient matching attack is finally performed on the denoised gradient. The present invention reveals the vulnerability of the noise-perturbed ViT model in the face of gradient leakage attacks, and provides an adversarial perspective and thinking inspiration for the construction of the ViT privacy protection mechanism. Summary of the Invention
[0005] This paper proposes a method for attacking noise-perturbed ViT gradient leakage in federated learning scenarios. This method involves proposing a noise prediction method based on a denoising diffusion model to recover the noise-perturbed gradient data. Based on this, a gradient leakage attack is then performed on the ViT model. This method enables attacks against protected gradient data.
[0006] The specific contents are as follows:
[0007] (1) This paper proposes a gradient feature extraction method based on ViT, aiming to provide a new technical path for the interpretability research of deep learning models by analyzing the gradient distribution characteristics during the ViT model training process. This method innovatively transforms gradient information into a structured feature representation, focusing on the extraction and analysis of positional encoding gradients to reveal the learning mechanism of the Transformer architecture in visual tasks.
[0008] In terms of method design, the present invention establishes a complete gradient feature extraction process, including three core steps: data preprocessing, model training, and gradient feature extraction. The gradient feature extraction stage adopts a hierarchical processing strategy to differentiate and extract conventional parameter gradients and position-encoded gradients, and uses dynamic resizing techniques to uniformly represent gradients of different dimensions as a standardized three-dimensional tensor. This method specifically designs a dedicated extraction mechanism for position-encoded gradients, effectively preserving the spatial structural characteristics of the gradient information by automatically calculating the optimal padding size and performing zero-padding operations.
[0009] The gradient processing flow is as follows:
[0010] 1) Flattening: Flatten all layer gradients into one-dimensional vectors;
[0011] 2) Size adjustment: Calculate the minimum square size g that meets the requirements;
[0012] 3) Zero value filling: Use F.pad to fill with zeros to make the vector length reach g 2 ;
[0013] 4) Reshape: Convert the padded gradient into a 3D tensor of 1×g×g.
[0014] Where: the calculation of the gradient filling size g adopts the rounding-up strategy, L is the original gradient vector length, δ is the adjustment factor (position encoding gradient δ = 2, conventional gradient δ = 0), and the calculation of the filling amount P is P = g 2 -L.
[0015] (2) This paper proposes a gradient denoising method based on diffusion probability modeling. By simulating the diffusion process of an image evolving from a clear state to a pure noise state, a neural network is constructed to learn the inverse denoising mapping, thereby predicting and removing image noise. This method is based on the Denoising Diffusion Probabilistic Models (DDPM). The core idea is to learn the process of reconstructing an image from noise during the training phase and to remove Gaussian noise from the gradient data during the generation phase.
[0016] like Figure 1As shown, the method proposed in the present invention includes the following three core stages: forward diffusion process modeling, reverse denoising process learning, and denoised image generation reasoning.
[0017] 1) Forward diffusion process modeling
[0018] Given an original image x0, a fixed Gaussian perturbation process is used to generate perturbed images x1, x2, ..., x at multiple time steps. T At each time step t∈{1,…,T}, the image is noised as follows:
[0019]
[0020] where β t is a predefined noise scheduling coefficient used to control the noise addition intensity.
[0021] 2) Reverse denoising process learning
[0022] Using neural network ε θ (x t ,t) learn the perturbation image x from any time step t Predict the original noise ∈ in , that is, build the following model:
[0023]
[0024] ε θ (x t ,t) is trained by minimizing the mean square error (MSE) loss function between the predicted noise and the true noise:
[0025]
[0026] 3) Denoised Image Generation Inference
[0027] First, the noise scale M is introduced. If the noise scale M is known, the scaling factor c = 1 / √(1+X 2 ), used to adjust the impact of noise. If M is unknown, randomly select c∈(0,1) as an empirical adjustment parameter.
[0028] Next, initialize the noise gradient X T′ :make That is, the input noise gradient is scaled to adapt it to the subsequent denoising process.
[0029] (3) This paper proposes an image privacy reconstruction attack method based on ViT gradient matching: through an optimization-based method, combined with model gradients and position-embedded sensitive information, it achieves accurate restoration of the input image. This method simulates an attacker who obtains model gradient information and then performs reverse optimization on the model input data to restore the original image content, revealing the potential privacy leakage risks of deep models during training or inference.
[0030] like Figure 2 As shown in the figure, the method mainly includes three core modules: label inference module, diffusion leakage (DL) loss construction module and virtual input optimization module.
[0031] 1) Label Inference Module: Based on the symbolic characteristics of the output layer gradient distribution, a heuristic strategy is used to determine the true label of the input image, thereby achieving label recovery and further improving the attack effect and targeting.
[0032] 2) Diffusion leakage loss construction module: Constructs a multi-component loss function for reconstructing and restoring the image, integrating gradient matching and position embedding constraints. The specific definitions are as follows:
[0033] Gradient matching term (L2 distance):
[0034] Position embedding gradient similarity (cosine similarity):
[0035] The total DA loss is:
[0036] Where α is a weight parameter, which is used to adjust the balance between the two loss components.
[0037] 3) Virtual Input Optimization Module: This module randomly initializes an input image tensor, dummy_x, and performs iterative backpropagation optimization on the total DL loss, so that the gradient generated by dummy_x gradually approaches the gradient of the real image, thereby gradually reconstructing the visual content of the input image in pixel space. Finally, the PSNR is used to evaluate the similarity between the restored image and the real image to measure the effectiveness of the attack. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a schematic diagram of the de-noising model training and reasoning process of the present invention;
[0039] Figure 2 Schematic diagram of the process of generating raw data images for the optimization-based method; DETAILED DESCRIPTION
[0040] The present invention is a method for attacking gradient leakage against ViT. The following uses the ViT model for image classification as an example to illustrate the implementation of the present invention. Those skilled in the art should understand that these implementations are intended only to illustrate the technical principles of the present invention and are not intended to limit the scope of application of the present invention.
[0041] The development language is Python, the deep learning framework is Pytorch, and the dataset used is CIFAR10. The specific steps are as follows:
[0042] Step 1: Create a training dataset. Input the CIFAR10 dataset into the ViT model for simulated federated learning. Use the autograd method to extract the gradients of the model parameters as the training dataset. Transform and pad the dataset to a suitable input for the diffusion model (in this case, 316 x 316). Finally, 1532 training set gradient data points are obtained.
[0043] Step 2: Diffusion model training. Use the linear_beta_schedule scheduler for diffusion model training, with base settings of beta_start = 0.00001 and beta_end = 0.002. Use the Adam optimizer with an initial learning rate of 0.0001. Set the batch size to 8, with each batch consisting of 16 matching pairs of samples from the training set described in Step 1. Ten training epochs, based on the 1,000-step diffusion step, yield good results.
[0044] Step 3: Gradient perturbation. Input the image into the ViT model and use the autograd method to extract the gradient of the model parameters. According to the privacy budget ε = 10, the failure probability δ tongya =10e-5, gradient clipping threshold dp_clip=1, learning rate lr=0.0001, batch size dp_batch=8 for noise addition, and calculate the privacy sensitivity and noise scale, and perform Gaussian noise perturbation on the extracted gradient data.
[0045] Step 4: Gradient denoising. The perturbed noise is resized to 316*316 using the above method, which is suitable for the diffusion model input. This is then fed into the diffusion model. After inference by the diffusion model, the gradient data is output after the noise is removed. This denoised gradient data is then restored to its original shape by reversing the above adjustment steps.
[0046] Step 5: Gradient matching attack. The denoised gradient is approximated as the original gradient, and the original data is restored from it. In this example, an optimization-based method is adopted, using the Adam optimizer, a random noise image as the optimization target, and lr = 0.5. The diffusion leakage loss value includes the gradient loss and ViT's unique position encoding loss, and they are weighted added. In this example, α = 1.5 is used. In each round of iteration, the noisy image is input into ViT to calculate the gradient, and the loss value is calculated with the denoised gradient. The Zongsheng image is optimized by this loss value. Experiments show that around iter = 400, the noisy image has been restored to a high level, and the PSNR value of the final image reaches more than 35, successfully realizing the gradient leakage attack of the noise-perturbed Vision Transformer model.
Claims
1. A gradient matching attack method for the noise-perturbed Vision Transformer model, which implements a gradient leakage attack on the Vision Transformer model with noise perturbations in a federated learning scenario. The method is characterized by: The diffusion model is used to train the denoising process of the gradient perturbed by Gaussian noise, and then the recovered approximate gradient is used to perform a gradient matching attack to infer and recover the training data. The purpose is to reveal the vulnerability of the noise-perturbed Vision Transformer model to gradient matching attacks. The method includes the following: S1. A gradient feature extraction method based on ViT aims to provide a new technical path for interpretability research of deep learning models by analyzing the gradient distribution characteristics during ViT model training. This method innovatively transforms gradient information into structured feature representations, focusing on the extraction and analysis of positional encoding gradients to reveal the learning mechanism of the Transformer architecture in vision tasks. S2. An image privacy reconstruction attack method based on ViT gradient matching, which achieves accurate restoration of the input image through an optimization-based method combined with model gradient and position-embedded sensitive information; this method simulates the attacker to reversely optimize the model input data when obtaining model gradient information, thereby restoring the original image content, revealing the potential privacy leakage risks of deep models during training or inference.
2. A gradient feature extraction method based on ViT according to claim 1, characterized in that, The method described fills and deforms the gradient to adapt to the input and output of the diffusion model, and performs a reverse operation after generation to restore it to the original gradient shape, and then performs subsequent operations. The specific content is as follows: The gradient feature extraction stage uses a hierarchical processing strategy to distinguish and extract conventional parameter gradients and position-encoded gradients. Dynamic resizing techniques are used to uniformly represent gradients of different dimensions as a standardized three-dimensional tensor. A dedicated extraction mechanism for position-encoded gradients is designed, which effectively preserves the spatial structural characteristics of the gradient information by automatically calculating the optimal padding size and performing zero-padding operations. The gradient processing process includes: 1) flattening: flattening all layer gradients into a one-dimensional vector; 2) resizing: calculating the minimum square size g that meets the conditions; 3) zero padding: using F.pad to perform zero padding to make the vector length reach g 2 4) Reshape: Convert the filled gradient into a 3D tensor of 1×g×g; where the gradient filling size g is calculated using a rounding-up strategy, L is the length of the original gradient vector, δ is the adjustment factor (position-encoded gradient δ=2, conventional gradient δ=0), and the filling amount P is calculated as P=g 2 -L.
3. The image privacy reconstruction attack method based on ViT gradient matching according to claim 1 is characterized in that: The method specifically includes the following steps: using a heuristic strategy to determine the true label of the input image based on the sign characteristics of the output layer gradient distribution, thereby achieving label recovery; introducing a diffusion leakage (DL) loss to guide the recovery from the virtual input image to the original image during the gradient matching attack process, which is specifically defined as follows: The gradient matching term (L2 distance): Position embedding gradient similarity (cosine similarity): α is the weight coefficient, which is used to adjust the weight of the position encoding gradient in the total loss; finally, the peak signal-to-noise ratio (PSNR) is used to evaluate the effect of the generated attack image.
Citation Information
Cited By
Federal learning parameter protection method and system based on reversible diffusion disturbance
CN121457568A