Robust deep learning model training method based on meta fine tuning adversarial training and related equipment
By introducing a dual-branch structure and meta-fine-tuning strategy into the deep learning model, dynamically updating the model parameter weights, the problem of insufficient robustness in the combat attacks in the existing technology is solved, efficient defense against multiple attack types is achieved, and computational costs are reduced.
Patent Information
- Application Number
- CN202510228170.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing deep learning models are not robust enough when facing multiple adversarial attacks, especially l∞ attacks and complex combination attacks, and the traditional adversarial training methods are costly, making it difficult to improve the defense capabilities of multiple attacks at the same time.
Adversarial training method based on meta-fine-tuning is adopted, by constructing L∞ adversarial training branches and combining adversarial training branches, the model parameter weights are dynamically updated, and the robustness between the two is balanced through the fusion strategy, and meta-fine-tuning is used to reduce computational costs using the pre-trained model.
It significantly improves the robustness of the model against l∞ attacks and compound attacks, reduces the computational cost, avoids the decline in robustness under multiple attack types, and enhances the generalization ability of the model.
Smart Images

Figure CN120163191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, specifically to the adversarial robustness of deep learning models, and particularly to a training method and related devices for a robust deep learning model based on meta-fine-tuning adversarial training. Background Art
[0002] Deep neural networks (DNNs) have performed excellently in many tasks in recent years, but their robustness is severely threatened by adversarial attacks. Adversarial attacks can deceive neural networks by adding tiny but targeted perturbations, causing them to misclassify input data. This vulnerability has led to extensive research on adversarial training (AT) methods aimed at improving the robustness of models under adversarial attacks.
[0003] Among existing defense methods, adversarial training (AT) is considered one of the most effective means of resisting adversarial attacks. AT (such as PGD-based adversarial training) significantly improves the robustness of models by generating adversarial samples and adding them to the training set, especially for attacks within a specific norm (such as l ∞ attack) and performs well. However, traditional AT methods have a large computational cost. Especially for large neural networks and large-scale datasets, the training time and resource consumption are extremely high. In addition, the robustness of traditional AT methods is usually limited to a single type of adversarial attack and cannot cope with complex combined attacks or other multi-norm attacks.
[0004] To address this problem, researchers have proposed different improvement methods. For example, Generalized Adversarial Training (GAT) is a defense strategy against Composite Adversarial Attacks (CAA). GAT makes the model able to defend against multiple different types of attacks by introducing a combination of multiple adversarial attacks (such as l ∞ attack and semantic attack). The evaluation of the GAT model on the CARBEN (Composite Adversarial Robustness Benchmark) benchmark shows that when using a random scheduling strategy and an optimized scheduling strategy for combined attacks, the model can effectively cope with complex perturbations. However, the GAT method has a significant drawback: although it performs well in defending combined attacks, its robustness against l ∞ attack often decreases. For example, when dealing with AA (AutoAttack) and l ∞ attack, the performance of GAT is much worse than that of a model optimized specifically for l ∞ attack. This is because during the training process, excessive focus on combined attacks leads to the l ∞The robustness is weakened. Summary of the Invention
[0005] In order to simultaneously improve the defense ability of deep learning models against multiple adversarial attacks and reduce the computational cost, the present invention proposes a training method and related devices for a robust deep learning model based on meta-fine-tuning adversarial training, which can take into account the robustness against l ∞ attacks and combined attacks, and effectively reduce the computational cost.
[0006] In a first aspect, the present invention provides a training method for a robust deep learning model based on meta-fine-tuning adversarial training, where the deep learning model is used for image processing, including:
[0007] Step 1: Construct an L ∞ adversarial training branch and a combined adversarial training branch; wherein, the model parameters of the L ∞ adversarial training branch and the combined adversarial training branch are initialized with the parameters of the pre-trained model;
[0008] Step 2: Construct L ∞ norm adversarial samples and combined attack adversarial samples;
[0009] Step 3: Use the L ∞ norm adversarial samples to update the model parameters of the L ∞ adversarial training branch by gradient; use the combined attack adversarial samples to update the model parameters of the combined adversarial training branch by gradient;
[0010] Step 4: Dynamically update the model parameter weights of the L ∞ adversarial training branch and the combined adversarial training branch, and fuse the model parameters of the L ∞ adversarial training branch and the combined adversarial training branch according to the updated model parameter weights;
[0011] Step 5: Update the hybrid model parameters with the fused model parameters;
[0012] Step 6: Determine whether the branch update condition is satisfied. If so, update the model parameters of the L ∞ adversarial training branch and the combined adversarial training branch with the updated hybrid model parameters; if not, return to Step 2;
[0013] Step 7: Determine whether the iteration stop condition is satisfied. If so, end the training; if not, return to Step 2.
[0014] Further, Step 2 specifically includes:
[0015] Obtain the original image dataset and the generated image dataset generated by the diffusion model;
[0016] Sample and mix from the original image dataset and the generated image dataset to obtain a mixed image dataset;
[0017] Use different perturbation strategies to perturb the mixed image dataset to generate L ∞ -norm adversarial samples and combined attack adversarial samples.
[0018] Furthermore, in step 3, when performing gradient update on the model parameters of the L ∞ adversarial training branch, the training objective function is:
[0019]
[0020] where θ1 represents the model parameters of the l ∞ adversarial training branch, represents the cross-entropy loss, represents the Kullback-Leibler divergence, λ is a hyperparameter, represents using l ∞ -PGD adversarial attack, ε represents the perturbation range of the PGD attack, x represents the clean sample, x adv is the perturbed adversarial sample, y represents the label corresponding to the clean sample, represents the data distribution of the clean sample, represents the data distribution of the adversarial sample, f represents the deep learning model, represents the expectation.
[0021] Furthermore, in step 3, when performing gradient update on the model parameters of the combined adversarial training branch, the training objective function is:
[0022]
[0023] where θ2 represents the model parameters of the combined adversarial training branch, Ω = {A1, A2, …, A N} is the attack set, E = {ε1, …, ε N} is the set of perturbation ranges of the corresponding attacks, represents the cross-entropy loss, f represents the deep learning model, x adv is the perturbed adversarial sample, y represents the label corresponding to the clean sample, represents the data distribution of the clean sample, represents the data distribution of the adversarial sample, represents the expectation.
[0024] Furthermore, performing gradient update on the model parameters of the L ∞ adversarial training branch includes:
[0025]
[0026] θ′1 ← τ·θ′1+(1 - τ)·θ1
[0027] Among them, γ represents the learning rate, τ represents the exponential decay rate, and θ′1 represents the updated model parameter.
[0028] Furthermore, gradient update is performed on the model parameters of the combined adversarial training branch, including:
[0029]
[0030] θ′2 ← τ·θ′2+(1 - τ)·θ2
[0031] Among them, γ represents the learning rate, τ represents the exponential decay rate, θ′2 represents the updated model parameter, and N represents the number of types of combined attacks.
[0032] Furthermore, in step 5, the hybrid model parameters are updated using the fused model parameters, including:
[0033] θ g ← τ·θ g +(1 - τ)·(β·θ′1+(1 - β)·θ′2)
[0034] Among them, τ represents the exponential decay rate, θ g represents the hybrid model parameter, θ′1 represents the model parameter of the L ∞ adversarial training branch after update, θ′2 represents the model parameter of the combined adversarial training branch after update, and β represents the weight of the model parameter of the branch.
[0035] In a second aspect, the present invention provides a training device for a robust deep learning model based on meta-fine-tuning adversarial training, including:
[0036] A dataset construction module, configured to obtain an original image dataset and a generated image dataset generated by a diffusion model;
[0037] A batch sampling module, configured to sample and mix the original image dataset and the generated image dataset to obtain a mixed image dataset;
[0038] A sample perturbation module, configured to generate L ∞ norm adversarial samples and combined attack adversarial samples based on the mixed image dataset;
[0039] A branch model parameter update module, configured to use the L ∞ norm adversarial samples for L ∞Perform gradient update on the model parameters of the adversarial training branch; use the combined attack adversarial samples to perform gradient update on the model parameters of the combined adversarial training branch;
[0040] A branch parameter fusion module for dynamically updating the L ∞ The model parameter weights of the adversarial training branch and the combined adversarial training branch, and perform fusion on the model parameters of the L ∞ Adversarial training branch and the combined adversarial training branch;
[0041] A hybrid model parameter update module for updating the hybrid model parameters using the fused model parameters;
[0042] A judgment module for judging whether the branch update condition is satisfied. If so, update the model parameters of the L ∞ Adversarial training branch and the model parameters of the combined adversarial training branch; and judge whether the iteration stop condition is satisfied. If so, end the training.
[0043] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0044] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0045] The beneficial effects of the present invention are:
[0046] (1) Defend against multiple adversarial attacks simultaneously
[0047] Traditional adversarial training methods, such as PGD-based l ∞ Adversarial training, although it can effectively defend against l ∞ Attacks, but performs poorly against complex combined attacks (such as color perturbation and geometric transformation). Although the GAT method performs well under combined attacks, it will weaken the model's robustness against l ∞ Attacks. Through the dual-branch structure design of the present invention, the model can not only maintain high robustness against l ∞ Attacks, but also enhance the defense ability against complex combined attacks. Compared with GAT, the present invention significantly improves the defense performance against combined attacks without sacrificing the defense effect against l ∞ Attacks.
[0048] (2) Improve the robustness against l ∞ Attacks while taking into account combined attacks
[0049] When the GAT method optimizes the combined attack defense, it often causes the model to be less robust against l ∞ attacks. Through the dynamic weight fusion strategy, the present invention balances the parameter updates of the l ∞ attack and the combined attack branches in each round of training, ensuring that the model maintains high robustness under both types of attacks. Compared with the GAT method, the present invention can better balance various attack defense requirements and solves the problem of l ∞ decreased robustness in existing methods.
[0050] (3) Significantly reduce the computational cost and reduce robust overfitting
[0051] The present invention adopts a meta-fine-tuning strategy based on a pre-trained model, uses an existing l ∞ robust model as the initialization model for fine-tuning, greatly reducing the training time and computational resource consumption. Compared with the training method from scratch, the fine-tuning strategy not only maintains the high robustness of the model but also significantly reduces the computational cost. In addition, by introducing the extended data generated by the diffusion model, the problem of robust overfitting of the model under specific attacks is effectively alleviated, and the generalization ability of the model is improved. Description of the Drawings
[0052] Figure 1 It is one of the schematic flowcharts of a training method for a robust deep learning model based on meta-fine-tuning adversarial training provided by an embodiment of the present invention;
[0053] Figure 2 It is another schematic flowchart of a training method for a robust deep learning model based on meta-fine-tuning adversarial training provided by an embodiment of the present invention;
[0054] Figure 3 It is the curve of the performance change with the number of training rounds under the single-branch structure provided by an embodiment of the present invention;
[0055] Figure 4 It is the curve of the robustness change with the number of training rounds under the double-branch structure provided by an embodiment of the present invention;
[0056] Figure 5 It is the schematic structural diagram of a training device for a robust deep learning model based on meta-fine-tuning adversarial training provided by an embodiment of the present invention;
[0057] Figure 6 It is the structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0058] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] The present invention adopts a dual-branch structure and a meta-learning strategy to improve the robustness of the model against l ∞ -norm adversarial attacks and enhance the model's ability to resist combined adversarial attacks. The present invention can be applied to deep learning tasks such as image classification and object detection that require robustness.
[0060] As Figure 1 shown, the embodiments of the present invention propose a training method for a robust deep learning model based on meta-fine-tuning adversarial training, which is used to improve the robustness of the model against both l ∞ -norm attacks and combined adversarial attacks. The method includes the following steps:
[0061] S101: Construct an L ∞ -adversarial training branch and a combined adversarial training branch, and initialize the model parameters of the L ∞ -adversarial training branch and the combined adversarial training branch using the parameters of the pre-trained model;
[0062] S102: Construct L ∞ -norm adversarial samples and combined attack adversarial samples;
[0063] S103: Use the L ∞ -norm adversarial samples to update the model parameters of the L ∞ -adversarial training branch by gradient; use the combined attack adversarial samples to update the model parameters of the combined adversarial training branch by gradient;
[0064] S104: Dynamically update the model parameter weights of the L ∞ -adversarial training branch and the combined adversarial training branch, and fuse the model parameters of the L ∞ -adversarial training branch and the combined adversarial training branch according to the updated model parameter weights;
[0065] S105: Update the hybrid model parameters using the fused model parameters;
[0066] S106: Determine whether the branch update condition is satisfied. If so, update the L ∞Update the model parameters of the adversarial training branch and the combined adversarial training branch; if not, return to step 1; for example, every c steps, update the model parameters of the two branches using the updated hybrid model parameters.
[0067] S107: Determine whether the iteration stop condition is satisfied. If so, end the training; if not, return to step 1; for example, set the maximum number of iterations, and when the maximum number of iterations is reached, end the training.
[0068] In the training method of the robust deep learning model provided by the embodiments of the present invention, it mainly includes a two-branch structure. One branch (hereinafter referred to as: l ∞ adversarial training branch or l ∞ attack-robust model M1) is used to defend against l ∞ attacks, and the other branch (hereinafter referred to as: CAA adversarial training branch or CAA attack-robust model M2) is used to defend against complex combined attacks. Initialize the model parameters of the two branches using a robust pre-training model, and during the training process, mix the model parameters of the two branches to obtain the parameters of the hybrid model M g and re-initialize the model parameters of the above two branches using the parameters of the hybrid model M g for the next iterative training. During the training process, for the l ∞ adversarial training branch, use l ∞ norm adversarial samples as input to train the model M1, and update the parameters updated in each batch to the hybrid model M g at a certain ratio; for the CAA adversarial training branch, apply attacks to clean images in random order from the attack set, and then use the attacked images as input to train the model M2, and update the parameters updated in each batch to the hybrid model M g at a certain ratio.
[0069] Task is a core concept in meta-learning. In the meta-fine-tuning adversarial training framework designed by the training method of the robust deep learning model provided by the embodiments of the present invention, it is divided into two branches, corresponding to two different task division methods. The l ∞ adversarial training branch trains on L ∞ adversarial samples, and the goal is to improve the L ∞ robustness; the CAA adversarial training branch trains on combined adversarial samples, and the goal is to improve the resistance to non-L pRobustness of novel adversarial attacks on norms. Among them, each branch updates the parameters of its respective branch based on the initial parameters provided by the hybrid model, which is equivalent to quickly adapting to and learning its respective tasks, corresponding to the process of fast adaptation in the inner loop of meta-learning; after the branches update for several steps based on the initial parameters, they perform weighted summation and then update the parameters of the hybrid model, corresponding to the outer loop process of meta-learning; after several steps, the two branch models are re-initialized with the parameters of the hybrid model, which is equivalent to starting the inner loop process of the next step. The present invention reduces the training cost and improves the generalization ability of the model by combining the pre-trained model and meta-learning technology.
[0070] In one embodiment, as Figure 2 shown, the samples participating in the training are a mixed image dataset formed by mixing the original image dataset and the generated image dataset generated by the diffusion model in a certain proportion η. Specifically, in each iteration process, a certain proportion of samples are sampled from the original image dataset, and a certain proportion of samples are sampled from the generated image dataset for mixing to form a batch of clean samples, and different strategies are used to perturb the batch of clean samples to construct the L ∞ norm adversarial samples and combined attack adversarial samples of this batch.
[0071] Specifically, in this embodiment, the CIFAR-10 dataset is used as the original image dataset, which contains 50,000 training pictures and 10,000 test pictures, and the image resolution is 32×32×3; this dataset is widely used in robustness research. The pre-trained weights of the class-conditional EDM diffusion model are used to generate 1 million and 5 million extended CIFAR-10 images respectively. For the generation of 5 million extended images, 500,000 images can be directly generated for each class, and their pseudo-labels are directly specified by the class condition; for the generation of 1 million images, the pre-trained WRN-28-10 model is used to score each image, and the top 20% of the images with the highest scores are selected for each class. These data are used to supplement the original training set through the proportion η.
[0072] In one embodiment, let the weight parameters generated by the two branch models be θ1∈Θ and θ2∈Θ respectively, the number of training rounds be T, and the parameter state trajectories corresponding to the model training process be and For any parameter θ∈Θ, the loss expectation of the hybrid model parameter θ g on the two branch test sets is constrained by the following formula with probability 1-δ:
[0073]
[0074] In formula (1), is from the distribution A series of loss functions sampled under a given data distribution Under R T is the regret of the two branch models on the corresponding loss functions. As an implementable approach, the sum of the gaps between each branch model parameter and the loss of the corresponding branch optimal parameter is used as R T , and the corresponding calculation formula is:
[0075]
[0076] Equation (1) shows that if the errors of both branches can be reduced simultaneously, then the error bound of the hybrid model can be narrowed, thereby enabling the hybrid model to perform sufficiently well on the tasks of both branches. Therefore, different strategies can be applied to the two branch models to ensure that the error of each branch is small enough.
[0077] In one embodiment, based on the training strategy of the original pre-trained model, when performing gradient update on the model parameters of the adversarial training branch, the training objective function is: ∞ When performing gradient update on the model parameters of the adversarial training branch, the training objective function is:
[0078]
[0079] where, θ1 represents the model parameters of the l ∞ adversarial training branch, represents the cross-entropy loss, represents the Kullback-Leibler divergence, λ is a hyperparameter, represents the use of l ∞ -PGD adversarial attack, ε represents the perturbation range of the PGD attack, x represents a clean sample, x adv is the perturbed adversarial sample, y represents the label corresponding to the clean sample, represents the data distribution of the clean sample, represents the data distribution of the adversarial sample, f represents the deep learning model, represents the expectation.
[0080] In one embodiment, considering from the two aspects of robust parameter initialization and training efficiency, in the combined adversarial training branch, less rounds of fine-tuning can be performed based on the currently best-performing robust model, and at the same time, the data generated by the diffusion model is used to increase the robust generalization ability of this branch. When performing gradient update on the model parameters of the combined adversarial training branch, the training objective function is:
[0081]
[0082] where, θ2 represents the model parameters of the combined adversarial training branch, Ω = {A1, A2,..., A N} is the set of attacks, E = {ε1,..., ε N} is the set of perturbation ranges of the corresponding attacks, represents the cross-entropy loss, f represents the deep learning model, x adv is the adversarial sample after perturbation, y represents the label corresponding to the clean sample, represents the data distribution of the clean sample, represents the data distribution of the adversarial sample, represents the expectation.
[0083] Specifically, denote Given the assignment function π: Then the combined adversarial sample can be expressed as:
[0084] x adv = A π(N) (A π(N-1) (…A π(1) (x)))(5)
[0085] A n (x) is expressed as the optimization problem of the perturbation δ n :
[0086]
[0087] For the solution of the optimal attack order π * is to solve the maximization problem of equation (4):
[0088]
[0089] Optimizing a single attack only through equations (6 - 7) cannot guarantee the optimization of other attacks in the entire attack sequence. By defining the doubly stochastic scheduling matrix where ∑ i z ij = ∑ j z ij = 1, and the surrogate adversarial sample where transforms the maximization problem of equation (7) into the iterative update of the scheduling matrix
[0090]
[0091] In the formula, is the Sinkhorn normalization, which ensures that the matrix obtained after each iteration is still doubly stochastic. Then, the Hungarian assignment algorithm is used to obtain the optimized attack order:
[0092]
[0093] Considering the training efficiency, the combined semantic adversarial sample attacks participating in the training of this branch adopt a random order, that is, the allocation function in Equation (7) is a random permutation of, and the optimized attack order π * is used for attacks during evaluation. Hue attacks, saturation attacks, brightness attacks, contrast attacks, rotation attacks, and PGD attacks are adopted in the attack set Ω. Similar to the PGD attack algorithm, the perturbation δ n of the semantic adversarial attack is solved by the following formula:
[0094]
[0095] The parameter initialization of the hybrid model is provided by the publicly released l ∞ robust model. After the two branches train the model according to their respective task distributions, it is necessary to mix the parameters of the model with a certain strategy to obtain a single model with both types of robustness. The model weights are averaged by averaging the model weights of different training rounds to obtain a flat loss landscape, which can effectively improve the robustness of the model; on the other hand, neural networks have the property of mode connectivity, which means that different local minima obtained by gradient descent are connected by a simple path in the parameter space, and similar loss values can be obtained on this path. Using the above properties, model weight averaging is adopted on the two branches respectively. That is:
[0096] θ′1←τ·θ′1+(1-τ)·θ1(11)
[0097] θ′2←τ·θ′2+(1-τ)·θ2(12)
[0098] where θ′1 represents the model parameters of the updated L ∞ adversarial training branch, and θ′2 represents the model parameters of the updated combined adversarial training branch.
[0099] Then the two-branch models explore the optimal solutions of the parameters on their respective data distributions. To prevent the deviation of each branch during the training process, after a certain interval c, the mode connectivity property is used on the hybrid model to linearly combine the model parameters of the two branches, and the model weights are averaged to smooth the model parameters. When the hybrid model M g is updated, it updates its own model in the way of exponential moving average, that is, it mixes its own historical model parameters and the model parameters of the two branches in a certain proportion:
[0100] θ g ←τ·θ g +(1-τ)·(β·θ′1+(1-β)·θ′2)(13)
[0101] where θ grepresents the hybrid model parameters, τ is the decay rate of the exponential moving average, and β represents the weight of the model parameters of the branch.
[0102] In the embodiments of the present invention, the idea of meta - learning is used to re - initialize the two branch models with the parameters of the hybrid model. The initialized parameters include the knowledge on both branches. On this basis, the two branch models can obtain the ability to quickly adapt to the data distribution of the branch after a small number of training rounds.
[0103] Based on the above - mentioned embodiments, correspondingly, for L ∞ Gradient update of the model parameters of the adversarial training branch includes:
[0104]
[0105] θ′1←τ·θ′1+(1 - τ)·θ1
[0106] Among them, γ represents the learning rate, τ represents the exponential decay rate, and θ′1 represents the updated model parameters.
[0107] Gradient update of the model parameters of the combined adversarial training branch includes:
[0108]
[0109] θ′2←τ·θ′2+(1 - τ)·θ2
[0110] Among them, γ represents the learning rate, τ represents the exponential decay rate, θ′2 represents the updated model parameters, and N represents the number of types of combined attacks.
[0111] In the embodiments of the present invention, the two branches have different loss functions and data distributions, and adopt optimization strategies specific to their respective tasks. Finally, through meta - fine - tuning adversarial training, the hybrid model has the L ∞ norm adversarial attack robustness while improving the robustness against combined adversarial attacks. The meta - fine - tuning adversarial training framework combines three strategies: fine - tuning, meta - learning, and adversarial training. By dividing the training process into two branches, optimizing for different types of adversarial attacks respectively, and using the idea of meta - learning to integrate the gradient information of different tasks, the fast adaptation ability and generalization ability of the model are improved.
[0112] As an implementable manner, the present invention also provides a specific algorithm description, as shown in Algorithm 1 below.
[0113]
[0114]
[0115] Based on the existing adversarial training methods, the present invention has made important innovations and improvements, aiming to enhance the robustness of the model under various types of adversarial attacks, especially in defending against l ∞ norm attacks and complex combined adversarial attacks. Through multiple improvements such as the design of a dual-branch structure, dynamic model fusion, and a training mechanism based on extended data, the present invention achieves more efficient adversarial defense performance. The specific manifestations are as follows:
[0116] (1) Design of the dual-branch structure
[0117] The present invention proposes an adversarial training method based on a dual-branch structure, enabling the model to simultaneously handle l ∞ norm attacks and combined adversarial attacks (CAA) during the training process. The dual-branch structure specifically includes:
[0118] l ∞ Adversarial training branch: This branch is mainly used to handle l ∞ norm attacks, ensuring that the model maintains high robustness performance in existing robustness evaluation benchmarks (such as RobustBench). By generating l ∞ norm adversarial samples in each round of training and optimizing the model in combination with the TRADES loss, its ability to defend against l ∞ attacks is enhanced.
[0119] Combined adversarial training branch: The other branch focuses on defending against combined adversarial attacks, combining multiple attack types (including hue, saturation, brightness, contrast, rotation attacks, etc.), enabling the model to adapt to complex and diverse attack patterns. By introducing multiple perturbations, this branch significantly improves the robustness performance of the model under the CARBEN benchmark.
[0120] The design of this dual-branch structure solves the limitation of traditional adversarial training methods that can only target a single attack type, while enhancing the model's defense ability against complex combined attacks.
[0121] (2) Dynamic model weight fusion strategy
[0122] After each round of training, the present invention uses a dynamic model weight fusion strategy to fuse the model parameters of the two branches in a weighted manner to ensure the robustness balance of the model on multiple tasks. This strategy includes:
[0123] Weight fusion mechanism: During the fusion process, the weights assigned to the two branch models are adjusted according to their performance during the training process, enabling the model to gradually enhance its defense ability against complex combined attacks while maintaining l ∞ robustness.
[0124] Exponential Moving Average (EMA) Strategy: In the weight fusion process, the EMA method is introduced to make the model parameters smoother, reduce training fluctuations, and further improve the stability and generalization of the model.
[0125] This weight fusion strategy ensures the robustness balance of different attack types while effectively avoiding the problem of damage to the l ∞ attack robustness in the GAT method when defending against combined attacks, and improves the comprehensive defense ability of the model.
[0126] (3) Meta-Fine-Tuning Strategy Based on Pre-Trained Model
[0127] Based on the existing pre-trained robust model, the present invention proposes an adversarial training strategy based on meta-fine-tuning, which greatly reduces the computational cost of training a large-scale adversarial robust model from scratch. The improvements are mainly reflected in:
[0128] Robust Fine-Tuning of Pre-Trained Model: The present invention directly uses the existing l ∞ robust model (such as the model pre-trained by the AT-EDM method) as the initialization model, avoiding the high cost of training from scratch. On this basis, the model is further fine-tuned to adapt to new adversarial attack tasks and improve its robustness.
[0129] Parameter Initialization and Optimization of Meta-Learning: The parameters of the hybrid model are optimized and initialized through the idea of meta-learning to ensure that the model can start training from a better initial state, thereby accelerating the adaptation process of the model to multiple attacks. Every few rounds, the hybrid model is used to re-initialize the two branch models to make them more adaptable to new tasks.
[0130] This meta-fine-tuning-based strategy not only significantly reduces training resources and time but also ensures the improvement of the model's robustness under multiple attack types.
[0131] To verify the effectiveness of the proposed solution of the present invention, the following experimental data is also provided.
[0132] (1) Experimental Settings
[0133] Evaluation Benchmark: Multiple attack modes such as AA (AutoAttack), SA (Semantic Attack), and CAA (Combined Attack) are adopted to evaluate the defense ability of the model. Among them, the perturbation range of AA is 8 / 255, and other parameters are automatically set in the attack algorithm. Both SA and CA adopt the optimized scheduling order, and the number of iterations is set to 5. The combined attack includes 4 types of attacks, namely CAA 3a 、CAA 3b 、CAA 3c 、CAA full 。CAA 3ais a combination of hue, saturation attack and PGD attack, CAA 3b is a combination of hue, rotation attack and PGD attack, CAA 3c is a combination of brightness, contrast attack and PGD attack.
[0134] Comparison model: Compared with the current optimal l ∞ adversarial training model and GAT model to evaluate robustness and generalization ability. Other models include: a standard model trained under the WRN34-10 architecture; the first-ranked GAT model was selected from the CARBEN leaderboard, and models with two scheduling strategies (random strategy and optimized strategy) under the WRN34-10 architecture were selected; models of AT-EDM trained under two architectures of WRN28-10 and WRN70-16 were selected; l ∞ The TRADES model trained by adversarial training; the AWP (Adversarial Weight Perturbation) model that doubles the perturbation of the input and weights during adversarial training under the WRN34-10 architecture.
[0135] Training settings: l ∞ In the adversarial training branch, the PGD-10 attack is adopted, with a perturbation range of 8 / 255, a step size of 2 / 255, and 10 iteration steps. In the combined adversarial training branch, the complete combined attack CAA full is adopted, that is, it includes five semantic perturbations and the PGD-10 attack, and the scheduling strategy adopts the random strategy. Among them, the five semantic perturbations include hue, saturation, rotation, brightness and contrast attacks, and their perturbation ranges are [-π,π], [0.7, 1.3], [-10°, 10°], [-0.2, 0.2], [0.7 - 1.3] respectively. The perturbation range of the PGD-10 attack is also 8 / 255, and the iteration steps of all single attack components are 10. The backbone network of the model adopts two structures of WRN-28-10 and WRN-70-16. The maximum number of training epochs is set to 10, the data batch size is set to 256, the proportion η of the generated data mixed in each batch of data is set to 0.5, the exponential moving average decay rate τ = 0.995, and the mixing coefficient β is dynamically adjusted with a piecewise linear function during training. The mixing coefficients in the 1st / 3rd / 6th / 10th training epochs are set to 1, 0.8, 0.6, and 0.4 respectively. Every c = 5 epochs, the two branch models are re-initialized with the mixed model, and the mixed model saves the optimal model according to the robust accuracy under PGD-40 and the combined attack.
[0136] (2) Experimental results
[0137] Experiment 1 compared the robustness performance of the method of the present invention and existing adversarial training methods under AA and CAA, as shown in Table 1.
[0138] Table 1 Defense performance of different models under attack combinations (unit: %)
[0139]
[0140]
[0141] As can be seen from Table 1, under different types of combined attacks, the method of the present invention consistently outperformed all other methods. Under the three types of combined attacks of CAA 3a , CAA 3b and CAA 3c , the method of the present invention exceeded the best method by more than 10%. For example, under the CAA 3a attack, compared with AT-EDM, it exceeded by 12%. Under CAA 3b , compared with GAT-fs, it exceeded by 12.3%. Under the CAA 3c attack, compared with GAT-fs, it exceeded by 14.6%. And under the full combined attack CAA full , it also exceeded the best model GAT-fs on the CARBEN leaderboard by 1.5%. In terms of clean accuracy and robust accuracy, the method of the present invention still maintained or even slightly exceeded the levels of normal models and original fine-tuned models. In terms of AA robust accuracy, compared with the original pre-trained model AT-EDM, the robust accuracy of the model of the present invention only decreased slightly, while for the GAT model compared with the original pre-trained model, its AA robust accuracy decreased by more than 10%. The current state-of-the-art AT-EDM for maintaining AA robustness has greatly reduced the impact of robust overfitting, but it cannot guarantee robustness against SA and CAA. For example, the accuracy for SA is 18.7, and the accuracy for the full combined attack is only 4.5%.
[0142] Experiment 2 compared the advantages of the multi-branch structure over the single-branch structure. Under the single-branch structure, the models of the adversarial training branch (branch 1) and the combined adversarial training branch (branch 2) of the present invention were respectively initialized with the pre-trained model of AT-EDM. During the training process, model fusion was not performed on the models of the two branches, and data generated by the diffusion model was used in each data batch, and they were trained for 10 epochs respectively. The change curves of clean accuracy, PGD attack accuracy, and combined attack accuracy (random strategy) during the training process were plotted, as ∞ shown. Figure 3 shown.
[0143] Among them, Clean1 and PGD1 represent l ∞Training curves of the adversarial training branches. Clean2, PGD2, and CAA2 represent the training curves of the combined adversarial training branches. From Figure 3 It can be seen that the clean accuracy and the PGD attack robust accuracy of the two branches change relatively smoothly. There is a gap of about 7% between Branch 1 and Branch 2 in terms of PGD robust accuracy, indicating that the addition of the SA component in Branch 2 has a certain impact on the robustness against the l ∞ attack. The change in the robustness of Branch 2 against the combined attack is relatively drastic. On the one hand, it is due to the trade-off between the l ∞ attack and SA. On the other hand, it is difficult for a single-branch model to stably learn new combined attack components within a relatively short number of training rounds.
[0144] The training curve of the hybrid model under the double-branch structure is as Figure 4 shown. The change trends of the clean accuracy and the PGD attack robust accuracy curves are the same as those of the curves under the single-branch structure, both changing relatively smoothly. The combined attack robust accuracy curve is significantly different from that under the single-branch structure, showing a trend of fluctuating upward. This indicates that the method of the present invention can stabilize the learning process of the model, enabling the learning of the combined attack components to maintain a relatively upward trend.
[0145] In terms of training time consumption, the method of the present invention only requires 10 training rounds. Training 10 rounds for the WRN28-10 architecture on a single NVIDIA RTX A100 GPU takes 5.2 hours, and training 10 rounds for the WRN70-16 architecture on three NVIDIA RTX A100 GPUs takes 17.3 hours. While GAT uses an ordinary Madry adversarial training model for fine-tuning and needs to train more than 100 rounds on multiple GPU cards to achieve better results.
[0146] Based on the same inventive concept, as Figure 5 shown, the present invention also provides a training device for a robust deep learning model based on meta fine-tuning adversarial training, including a dataset construction module, a batch sampling module, a sample perturbation module, a branch model parameter update module, a branch parameter fusion module, a hybrid model parameter update module, and a judgment module.
[0147] Among them, the dataset construction module is used to obtain the original image dataset and the generated image dataset generated by the diffusion model; the batch sampling module is used to sample and mix the original image dataset and the generated image dataset to obtain a mixed image dataset; the sample perturbation module is used to generate L ∞ norm adversarial samples and combined attack adversarial samples based on the mixed image dataset; the branch model parameter update module is used to use the L ∞ norm adversarial samples to the L ∞Perform gradient update on the model parameters of the adversarial training branch; use the combined attack adversarial samples to perform gradient update on the model parameters of the combined adversarial training branch; the branch parameter fusion module is used to dynamically update the ∞ model parameter weights of the adversarial training branch and the combined adversarial training branch, and based on the updated model parameter weights, the ∞ model parameters of the adversarial training branch and the combined adversarial training branch are fused; the hybrid model parameter update module is used to update the hybrid model parameters with the fused model parameters; the judgment module is used to judge whether the branch update condition is satisfied, and if so, use the updated hybrid model parameters to ∞ update the model parameters of the adversarial training branch and the model parameters of the combined adversarial training branch; and judge whether the iteration stop condition is satisfied, and if so, end the training.
[0148] The deep learning model trained by the training device of the robust deep learning model provided by the embodiments of the present invention can significantly improve the defense performance against combined attacks without sacrificing the ∞ attack defense effect.
[0149] Figure 6 Illustrates a schematic physical structure diagram of an electronic device, as Figure 6 shown. The electronic device may include: a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logical instructions in the memory 603 to execute the training method of the robust deep learning model. The method includes: Step 1: Construct ∞ an adversarial training branch and a combined adversarial training branch; wherein, use the parameters of the pre-trained model to ∞ initialize the model parameters of the adversarial training branch and the combined adversarial training branch; Step 2: Construct ∞ norm adversarial samples and combined attack adversarial samples; Step 3: Use the ∞ norm adversarial samples to perform gradient update on the model parameters of the ∞ adversarial training branch; use the combined attack adversarial samples to perform gradient update on the model parameters of the combined adversarial training branch; Step 4: Dynamically update the ∞ model parameter weights of the adversarial training branch and the combined adversarial training branch, and based on the updated model parameter weights, the ∞Fuse the model parameters of the adversarial training branch and the combined adversarial training branch; Step 5: Update the hybrid model parameters with the fused model parameters; Step 6: Determine whether the branch update condition is satisfied. If so, update the L with the updated hybrid model parameters ∞ Update the model parameters of the adversarial training branch and the combined adversarial training branch; If not, return to Step 1; Step 7: Determine whether the iteration stop condition is satisfied. If so, end the training. If not, return to Step 1.
[0150] In addition, when the logical instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0151] The embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the training method of the robust deep learning model provided by each of the above method embodiments.
[0152] The embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the robust deep learning model provided by each of the above method embodiments.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for training a robust deep learning model based on meta-fine-tuning adversarial training, wherein the deep learning model is used for image processing, characterized in that: include: Step 1: Build L ∞ adversarial training branch and combined adversarial training branch; wherein the parameters of the pre-trained model are used to train the L ∞ Initialize the model parameters of the adversarial training branch and the combined adversarial training branch; Step 2: Build L ∞ Norm adversarial samples and combined attack adversarial samples; Step 3: Use the L ∞ Norm adversarial examples for L ∞ Performing gradient updates on the model parameters of the adversarial training branch; performing gradient updates on the model parameters of the combined adversarial training branch using the combined attack adversarial sample; Step 4: Dynamically update the L ∞ The model parameter weights of the adversarial training branch and the combined adversarial training branch, and the L ∞ The adversarial training branch and the combined adversarial training branch are combined with model parameters of the adversarial training branch; Step 5: Use the fused model parameters to update the hybrid model parameters; Step 6: Determine whether the branch update condition is met. If so, use the updated hybrid model parameters to update the L ∞ The adversarial training branch and the combination update the model parameters of the adversarial training branch; if not, return to step 2; Step 7: Determine whether the iteration stop condition is met. If so, end the training; if not, return to step 2.
2. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 1, characterized in that: Step 2 specifically includes: Obtaining an original image dataset and a generated image dataset generated by using a diffusion model; Sampling and mixing the original image dataset and the generated image dataset to obtain a mixed image dataset; Different perturbation strategies are used to perturb the mixed image dataset to generate L ∞ Norm adversarial examples and combined attack adversarial examples.
3. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 1, characterized in that: In step 3, ∞ When the gradient of the model parameters of the adversarial training branch is updated, the objective function of the training is: Among them, θ1 represents l ∞ Model parameters of the adversarial training branch, e CE represents the cross entropy loss, l KL represents the Kullback-Leibler divergence, λ is a hyperparameter, Indicates the use of l ∞ -PGD adversarial attack, ε represents the perturbation range of PGD attack, x represents the clean sample, x adv is the adversarial sample after perturbation, y represents the label corresponding to the clean sample, represents the data distribution of clean samples, represents the data distribution of adversarial samples, f represents the deep learning model, Express expectations.
4. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 1, characterized in that: In step 3, when the gradient of the model parameters of the combined adversarial training branch is updated, the objective function of the training is: Among them, θ2 represents the model parameters of the combined adversarial training branch, Ω={A1,A2,...,A N } is the attack set, E={ε1,...,ε N } is the set of disturbance ranges of the corresponding attacks, l CE represents the cross entropy loss, f represents the deep learning model, x adv is the adversarial sample after perturbation, y represents the label corresponding to the clean sample, represents the data distribution of clean samples, represents the data distribution of adversarial samples, Express expectations.
5. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 3, characterized in that: For L ∞ Perform gradient updates on the model parameters of the adversarial training branch, including: θ′1←τ·θ′1+(1-τ)·θ1 Among them, γ represents the learning rate, τ represents the exponential decay rate, and θ′1 represents the updated model parameters.
6. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 4, characterized in that: Perform gradient updates on the model parameters of the combined adversarial training branch, including: θ′2←τ·θ′2+(1-τ)·θ2 Among them, γ represents the learning rate, τ represents the exponential decay rate, θ′2 represents the updated model parameters, and N represents the number of combined attack types.
7. The method for training a robust deep learning model based on meta-fine-tuning adversarial training according to claim 1, characterized in that: In step 5, the hybrid model parameters are updated using the fused model parameters, including: i g ←t·i g +(1-τ)·(β·θ′1+(1-β)·θ′2) Where τ represents the exponential decay rate, θ g represents the mixed model parameters, θ′1 represents the updated L ∞ The model parameters of the adversarial training branch, θ′2 represents the model parameters of the updated combined adversarial training branch, and β represents the weight of the model parameters of the branch.
8. A training device for a robust deep learning model based on meta-fine-tuning adversarial training, characterized in that: include: A dataset construction module, used to obtain an original image dataset and a generated image dataset generated by a diffusion model; A batch sampling module, used for sampling and mixing the original image dataset and the generated image dataset to obtain a mixed image dataset; A sample perturbation module is used to generate L ∞ Norm adversarial samples and combined attack adversarial samples; The branch model parameter updating module is used to update the L ∞ Norm adversarial examples for L ∞ Performing gradient updates on the model parameters of the adversarial training branch; performing gradient updates on the model parameters of the combined adversarial training branch using the combined attack adversarial sample; Branch parameter fusion module, used to dynamically update the L ∞ The model parameter weights of the adversarial training branch and the combined adversarial training branch, and the L ∞ The adversarial training branch and the combined adversarial training branch are combined with model parameters of the adversarial training branch; A hybrid model parameter updating module, used to update the hybrid model parameters using the fused model parameters; The judgment module is used to judge whether the branch update condition is met. If so, the updated hybrid model parameters are used to update the L ∞ The model parameters of the adversarial training branch and the model parameters of the combined adversarial training branch are updated; and it is determined whether the iteration stop condition is met, and if so, the training is terminated.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Defense method for robust enhanced intelligent system based on diffusion model and adversarial training
CN121527480A