Unknown domain confrontation attack method based on gradient alignment
By introducing gradient alignment technology into the adversarial attack method in unknown domains, the problem of the reduced effectiveness of existing migration adversarial attacks in cross-scenarios is solved, and the ability to generate effective adversarial samples in unseen domains is realized, which significantly improves the success rate of cross-domain migration attacks.
Patent Information
- Application Number
- CN202510054988.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The effectiveness of existing migration adversarial attacks is significantly reduced across scenarios, making it difficult to generate effective adversarial samples in unseen domains.
A method of unknown domain adversarial attack based on gradient alignment is proposed. By searching for domain sensitive directions, the gradient of the alternative model is smoothed and aligned with the gradient distributed on the domain migration trajectory, thereby enhancing the effectiveness of cross-domain migration attacks.
This method can generate effective adversarial samples in unknown domains, significantly improving the success rate of cross-domain migration attacks and enhancing the reliability of adversarial samples.
Smart Images

Figure CN119990249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an unknown domain counterattack method based on gradient alignment, and belongs to the technical field of artificial intelligence information security. Background Art
[0002] In artificial intelligence neural networks, adversarial examples have raised many security issues. Studying adversarial examples of neural networks is of great significance to improving the reliability of artificial intelligence models in applications. However, the effectiveness of existing transfer adversarial attacks in cross-scenarios is significantly reduced. To address this phenomenon, the present invention proposes a domain-transferable adversarial attack (DTAA), which enables the surrogate model to generate effective adversarial examples even on previously unseen domains. Specifically, the method first searches for domain-sensitive directions and smoothes the gradients of the surrogate model along these directions. The present invention provides a theoretical analysis that illustrates how the proposed DTAA method aligns the gradients of the surrogate model with the gradients of the data distribution on the domain transfer trajectory. This alignment makes the gradients of the surrogate model closer to the gradients of the target model trained on the unknown domain, thereby enhancing the effectiveness of cross-domain transfer attacks. In addition, the present invention conducts extensive experiments using mature benchmarks on classification and verification tasks, and the results demonstrate the superiority of DTAA in cross-domain transfer attacks. Summary of the invention
[0003] In order to solve the problem that the effectiveness of existing migration adversarial attacks in cross-scenario situations is significantly reduced, the present invention proposes an unknown domain adversarial attack method based on gradient alignment to solve the above problem, which specifically includes:
[0004] Step 1: Obtain the source domain dataset D s ;
[0005] Step 2: In the source domain dataset D s Use on Training Methods Train Agent Models
[0006] Step 3: Generate transferable effective adversarial samples based on the proxy model, adversarial in the target dataset D t The target model for training Complete adversarial attacks on unknown domains.
[0007] Preferably, step 2 specifically includes:
[0008] Step 2.1: Model the adversarial perturbation generation method for the unknown domain and calculate the training method score;
[0009] Step 2.2: Introduce the transformation function T(·) to the source domain data sample x s Processing is performed to obtain cross-domain data representation T(xs ), define a neighborhood Ω with a radius of γ outside the source domain, and obtain the sample with the largest domain difference in the neighborhood within the domain
[0010] Step 2.3: Sampling iterative gradient ascent method approximates T(x s ), design the alignment loss function and classification loss function in the gradient neighborhood transformation process of the proxy model;
[0011] Step 2.4: Construct a total loss function based on the alignment loss function and the classification loss function to complete the proxy model training;
[0012] Training Methods The score expression is:
[0013]
[0014] In formulas (1) and (2), For training methods , A(·) is the backend of the transfer attack.
[0015] Preferably, step 2.2 specifically includes:
[0016] Step 2.2.1: Calculate the model gradients between different domains and introduce the transformation function T(·) for the sample x of the source domain dataset s Processing is performed to obtain cross-domain data representation T(x s );
[0017] Step 2.2.2: Define a neighborhood Ω with a radius of γ outside the source domain, based on the source model f θ The feature extraction component e(·) in s Mapping to feature space;
[0018] Step 2.2.3: Use the Frobenius norm || e(x i )-e(x j )|| F Calculate sample x i The features and samples x j The distance γ between the features, where x i ∈x s , x i ∈x s ;
[0019] Step 2.4: When γ is less than the preset value, maximize the distance ||e(x i )-e(x j )|| F , get the sample with the largest domain difference in the neighborhood Ω
[0020] The expression for aligning model gradients between different fields is:
[0021]
[0022] In formula (3), s1, s2…s l is the source domain, is the confidence of the proxy model for category y, F is the Frobenius norm, L is the source domain set, is a differential operator used to calculate the gradient;
[0023] The calculation formula for the sample with the largest domain difference in the neighborhood Ω is:
[0024]
[0025] Preferably, step 2.3 specifically includes:
[0026] Step 2.3.1: Use the iterative gradient ascent method to make the samples in the neighborhood Ω Approximate T(x s );
[0027] Step 2.3.2: After n iterations, the sample is obtained Approximate value of
[0028] Step 2.3.3: Design the alignment loss function L align Ensure that the gradients of the proxy model are aligned during the domain conversion process;
[0029] Step 2.3.4: Define the classification loss function L class Prevent the proxy model from overfitting on the source domain;
[0030] The expression of iterative gradient ascent method is:
[0031]
[0032] In formula (6), p represents the disturbance, which is defined as α is the balance coefficient, Indicates x s The neighborhood projection with the center and radius γ;
[0033] Alignment loss function L align The expression is:
[0034]
[0035] Classification loss function L class The expression is:
[0036]
[0037] In formulas (8) and (9), is a regularization term to enhance the smoothness of the model, μ is the balance coefficient, and CE() is the cross entropy loss function.
[0038] Preferably, the expression of the total loss function in step 2.4 is:
[0039] L=L align +η·L class (10);
[0040] In formula (10), η is the balance coefficient.
[0041] The beneficial effects of the present invention are:
[0042] 1. The present invention uses multiple backend attack methods to generate adversarial samples on the proxy model, which allows the method proposed in the present invention to enhance the existing transfer attack and make it suitable for cross-domain transfer attacks.
[0043] 2. The present invention attributes the key to successful transfer attacks to reducing the likelihood of adversarial samples for the data distribution obeyed by the training dataset. By reducing this likelihood, these transfer attacks can generate effective adversarial samples that have the potential to deceive any unknown target model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A flow chart of the unknown domain counterattack method based on gradient alignment provided by the present invention;
[0045] Figure 2 A flow chart of the gradient alignment method provided by the present invention;
[0046] Figure 3 A schematic diagram showing the comparison results of the DTAA method proposed in the present invention and other adversarial attack methods in enhancing the proxy model in cross-domain migration attacks;
[0047] Figure 4 A schematic diagram showing the comparison results of the gradient-based backend attack provided by the present invention combined with DTAA in the cross-domain migration attack;
[0048] Figure 5 A schematic diagram showing the comparison results of the transformation-based backend attack provided by the present invention combined with DTAA in the cross-domain migration attack;
[0049] Figure 6 A schematic diagram of the comparison results of the backend attack based on the objective function provided by the present invention combined with DTAA in the cross-domain migration attack. DETAILED DESCRIPTION
[0050] Specific implementation method 1: Combination Figure 1 and Figure 2 This embodiment is described. In this embodiment, the DTAA method is an unknown domain anti-attack method based on gradient alignment proposed by the present invention. Figure 1 and Figure 2 As shown, the steps of the unknown domain anti-attack method based on gradient alignment described in this embodiment include:
[0051] S1: We formulate adversarial perturbation generation for unknown domains as an optimization problem.
[0052] S101: Modeling adversarial perturbation generation methods for typical unknown domains;
[0053] In a cross-domain transfer attack, the attacker cannot access the target dataset D t It is important to have a dataset D in the source domain s A proxy model is trained on the target domain. In the test phase, The key challenge is to determine a suitable training method, denoted as Make The trained model Able to generate effective adversarial samples to counter the target domain D t The target model trained on Further, the question is how to find the optimal The score is defined as:
[0054]
[0055] In formulas (1) and (2), For training methods , and A(·) is the backend of the transfer attack.
[0056] S102: Formal definition of adversarial attack methods
[0057] S10201: A direct It is to align the model gradients between different fields, which can be expressed as:
[0058]
[0059] In formula (3), s1, s2…s l is the source domain, is the confidence of the proxy model for category y, F is the Frobenius norm, L is the source domain set, is a differential operator used to calculate the gradient; however, this approach is impractical. First, it is difficult to identify semantically aligned sample pairs across various domains. i ,x j ) are not semantically aligned, their different semantic information will interfere with the gradient alignment of proxy models between different domains. In addition, the source domain datasets are usually insufficient. All data may come from a single source domain, or the domain labels of the source domain are missing.
[0060] S10202: In order to promote cross-domain gradient alignment, this embodiment introduces a transformation function T(·) for the sample x in the source domain s Processing is performed to obtain cross-domain data representation T(x s ), the goal is to maximize T(x) while preserving semantic information s ) and x s In order to preserve x s The semantic information of the sample is obtained by defining a neighborhood Ω with a radius of γ. In this neighborhood, the sample is found It is related to x s Has the greatest field differences.
[0061] S10203: Source Model f θ It can be divided into feature extraction component e(·) and classification component c(·). For the source domain D s A sample x in s , e(·) will s Mapped to the feature space, this implementation uses the Frobenius norm || e(x i )-e(x j )|| F To quantize the sample x i and x j Feature e(x i ) and e(x j ). This distance originates from the difference between semantic information and domain information. When γ is small enough, x s and The semantic information difference between them becomes insignificant, allowing this implementation to focus on domain differences. By maximizing the distance || e(x i )-e(x j )|| F , this embodiment can As the sample with the largest domain difference in its neighborhood, it is defined as follows:
[0062]
[0063] S10204: This embodiment uses an iterative gradient ascent method to approximate T(x s ), the formula is as follows:
[0064]
[0065] In formula (6), p represents the disturbance, which is defined as α is the balance coefficient, Indicates x s The neighborhood projection with the center and radius γ;
[0066] S10204: In order to ensure that the gradients of the proxy model are aligned as much as possible during the domain conversion process, this embodiment designs an alignment loss L align ;
[0067] Alignment loss function L align The expression is:
[0068]
[0069] S10205: In order to prevent the proxy model from overfitting on the source domain, this implementation defines the classification loss:
[0070] Classification loss function L class The expression is:
[0071]
[0072] In formulas (8) and (9), is a regularization term to enhance the smoothness of the model, μ is the balance coefficient, and CE() is the cross entropy loss function.
[0073] S10206: Construct a total loss function based on the alignment loss function and the classification loss function to complete the proxy model training;
[0074] The expression of the total loss function is:
[0075] L=L align +η·L class (10);
[0076] In formula (10), η is the balance coefficient.
[0077] Once the proxy model is well trained, multiple backend attack methods can be used to generate adversarial samples on the proxy model, which allows the method of this embodiment to enhance existing transfer attacks and make them suitable for cross-domain transfer attacks.
[0078] S2: Alternative Model Domain Transfer Training Algorithm
[0079] In order to solve the optimization problem described in S1, this embodiment proposes an alternative model domain transfer training algorithm, the detailed process of the algorithm is as follows:
[0080]
[0081]
[0082] Specific implementation method 2: Combination Figure 3-6 This embodiment is explained. In order to verify the technical effect of the DTAA method proposed in the specific embodiment, the present invention is implemented and experimented on two publicly available datasets: MNIST-M and SVHN, with MNIST-M as the source and SVHN as the target, where MNIST-M is a color handwritten digital dataset and SVHN is a street photography digital dataset. In order to comprehensively evaluate the attack efficiency of the proposed unknown domain adversarial attack method on various recognition models, this embodiment uses VGG, ResNet, SENet and DenseNet models as source and target models for experiments. This embodiment compares the proposed DTAA method with other adversarial attack methods, and the results are as follows. Figure 3 shown.
[0083] Backend attacks can be divided into three main types: gradient-based, transformation-based, and target-based. In order to evaluate their effectiveness in cross-domain scenarios and study the possible performance improvement by combining with DTAA, this implementation tested multiple attacks of each type. The results are shown in the following table. Figure 4 , Figure 5 and Figure 6 As shown in Figure 2, although these methods have a high success rate in cross-model transfer attacks, their effectiveness is significantly reduced in cross-domain scenarios. However, by combining these backend attacks with DTAA, the present embodiment observed a significant improvement in the attack success rate (ASR). Among them, the STM backend performed outstandingly, with an average success rate of 77.27%, compared to Figure 4-6 It can be seen that the performance of these attack methods exceeds that of PGD, which indicates that the DTAA proposed in this embodiment can be effectively combined with these methods to release its model migration capability in the context of cross-domain migration attacks.
[0084] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.
Claims
1. An unknown domain adversarial attack method based on gradient alignment, characterized in that: The steps of the unknown domain adversarial attack method based on gradient alignment include: Step 1: Obtain the source domain dataset D s ; Step 2: In the source domain dataset D s Use on Training Methods Train Agent Models Step 3: Generate transferable effective adversarial samples based on the proxy model, adversarial in the target dataset D t The target model for training Complete adversarial attacks on unknown domains.
2. The unknown domain adversarial attack method based on gradient alignment according to claim 1 is characterized in that: Step 2 specifically includes: Step 2.1: Model the adversarial perturbation generation method for the unknown domain and calculate the training method score; Step 2.2: Introduce the transformation function T(·) to the source domain data sample x s Processing is performed to obtain cross-domain data representation T(x s ), define a neighborhood Ω with a radius of γ outside the source domain, and obtain the sample with the largest domain difference in the neighborhood within the domain Step 2.3: Sampling iterative gradient ascent method approximates T(x s ), design the alignment loss function and classification loss function in the gradient neighborhood transformation process of the proxy model; Step 2.4: Construct a total loss function based on the alignment loss function and the classification loss function to complete the proxy model training; Training Methods The expression is: In formula (1), For training methods score, A(·) is the backend of the transfer attack, and x is the target dataset D t The horizontal coordinate of the data in the graph, y is the target data set D t The vertical coordinate of the data; Training Methods The score expression is: In formula (2), x is the source domain dataset D s The horizontal coordinate of the data in the source domain is y, and y is the source domain dataset D s The vertical coordinate of the data.
3. The unknown domain adversarial attack method based on gradient alignment according to claim 2 is characterized in that: Step 2.2 specifically includes: Step 2.2.1: Calculate the model gradients between different domains and introduce the transformation function T(·) for the sample x of the source domain dataset s Processing is performed to obtain cross-domain data representation T(x s ); Step 2.2.2: Define a neighborhood Ω with a radius of γ outside the source domain, based on the source model f θ The feature extraction component e(·) in s Mapping to feature space; Step 2.2.3: Use the Frobenius norm || e(x i )-e(x j )|| F Calculate sample x i The features and samples x j The distance γ between the features, where x i ∈x s , x j ∈x s ; Step 2.4: When γ is less than the preset value, maximize the distance ||e(x i )-e(x j )|| F , get the sample with the largest domain difference in the neighborhood Ω The expression for aligning model gradients between different fields is: In formula (3), s1, s2…s l is the source domain, is the confidence of the proxy model for category y, F is the Frobenius norm, L is the source domain set, is a differential operator used to calculate the gradient; The calculation formula for the sample with the largest domain difference in the neighborhood Ω is:
4. The unknown domain adversarial attack method based on gradient alignment according to claim 2 is characterized in that: Step 2.3 specifically includes: Step 2.3.1: Use the iterative gradient ascent method to make the samples in the neighborhood Ω Approximate T(x s ); Step 2.3.2: After n iterations, the sample is obtained Approximate value of Step 2.3.3: Design the alignment loss function L align Ensure that the gradients of the proxy model are aligned during the domain conversion process; Step 2.3.4: Define the classification loss function L class Prevent the proxy model from overfitting on the source domain; The expression of iterative gradient ascent method is: In formula (6), p is the disturbance, defined as α is the balance coefficient, Indicates x s The neighborhood projection with the center and radius γ; Alignment loss function L align The expression is: Classification loss function L class The expression is: In formulas (8) and (9), is a regularization term to enhance the smoothness of the model, μ is the balance coefficient, and CE() is the cross entropy loss function.
5. The unknown domain adversarial attack method based on gradient alignment according to claim 2, characterized in that: The expression of the total loss function in step 2.4 is: L=L align +η·L class (10); In formula (10), η is the balance coefficient.