Adversarial sample generation method for destroying middle layer features
By perturbing features multiple times in the feature space along the direction of feature importance and introducing momentum to stabilize the update direction, the problem of low perturbation efficiency of existing feature-level attack feature is solved, and the generated adversarial samples have stronger migration.
Patent Information
- Application Number
- CN202510619864.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the prior art, the characteristic perturbation efficiency of feature-level attacks is low, resulting in insufficient migration of the generated adversarial samples.
By perturbing the features multiple times in the feature space along the direction of feature importance and introducing momentum to stabilize the update direction, the perturbed features are finally inverted to the pixel space to generate an adversarial sample.
The efficiency of feature perturbation is improved, the generated adversarial samples have stronger migration, and can more effectively mislead the intermediate layer features of different DNN models.
Smart Images

Figure CN120124709A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and particularly relates to a method for generating adversarial samples that disrupt intermediate layer features. Background Art
[0002] Deep neural networks (DNNs) have achieved great success in various machine learning tasks. However, a large amount of work has shown that DNNs are vulnerable to adversarial samples, which add carefully designed and imperceptible perturbations to clean images to mislead DNNs. The existence of adversarial samples has raised concerns about security-sensitive applications (such as autonomous driving and face recognition). The study of adversarial samples can not only help researchers understand the principles of adversarial samples and the deficiencies of DNNs, but also improve the defense capabilities of DNNs and the robustness of images, enabling them to remain stable and accurate in various application scenarios.
[0003] Many methods have been proposed to generate adversarial samples. According to the information of the target model available to the attacker, adversarial attacks can generally be divided into two categories: white-box attacks and black-box attacks. White-box attacks mean that the attacker can obtain all the knowledge of the target model (such as structure and parameters) to generate adversarial samples through gradient information. In contrast, in black-box attacks, the attacker cannot access the structure and parameters of the model, and thus cannot obtain gradient information, which makes black-box attacks more challenging and more in line with the actual situation.
[0004] Black-box attacks can be further divided into: transferability-based attacks and query-based attacks according to different attack strategies. Query-based attacks approximate gradient information through queries or use intelligent search algorithms to explore the input space to generate adversarial samples. However, query-based attacks are infeasible in many practical scenarios, such as face recognition and autonomous driving, because a large number of queries are not allowed in these scenarios. In contrast, transferability-based black-box attacks are more realistic and flexible because they do not require any knowledge of the target model. It first attacks a local white-box proxy model, and then directly transfers the obtained adversarial samples to an unknown target model. This cross-model attack ability of adversarial samples is called transferability.
[0005] Some studies have attempted to perform attacks at the intermediate layer to enhance the transferability of adversarial samples. These feature-level attacks do not directly interfere with the output layer of the proxy model, but achieve higher transferability by maximizing internal feature distortion. Since the most critical features are shared among different DNN models, feature-level attacks show promise in generating stronger transferable adversarial samples. Summary of the Invention
[0006] Objective of the Invention: The objective of the present invention is to solve the deficiencies existing in the prior art and provide an adversarial sample generation method for destroying intermediate layer features. The features are perturbed multiple times along the direction of feature importance in the feature space, and then the perturbed features are inverted onto the image to generate more transferable adversarial samples. The present invention can improve the efficiency of feature perturbation and generate more transferable adversarial samples.
[0007] Technical Solution: An adversarial sample generation method for destroying intermediate layer features of the present invention sets the classification model as , and for this classification model, the adversarial sample generation method P2FA is used to perturb the intermediate layer features of the classification model along the direction with a step size of to obtain the perturbed intermediate layer features ;
[0008] Among them, and respectively represent the input original image and the corresponding true label; represents the feature map of the th intermediate layer;
[0009] When performing the above perturbation T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into local optima. The specific process is as follows:
[0010] ;
[0011] ; Among them, is the decay factor in the momentum, is the perturbation step size, is the update direction of the intermediate layer features at the th iteration, || || 2 represents the L2 norm; is the feature importance, and the iteration formula of is as follows: ; ; n is the aggregation times for obtaining the feature importance, and the value range is 1 - N, represents the fitted image at the nth iteration, represents the perturbation magnitude, represents the cross-entropy loss; Finally, the perturbed features are inverted into the pixel space to obtain the corresponding adversarial sample , that is, inverting the features is summarized as the following optimization problem: ; Among them, represents the square of the L2 norm.
[0012] Furthermore, before performing T perturbations on the intermediate layer features of the classification model, first initialize the adversarial sample and the intermediate layer feature update direction , and the formulas are as follows:
[0013] ; ; ;
[0014] Then, during the above T perturbations, the adversarial sample and the intermediate layer feature update direction generated at the th iteration are respectively denoted as , ,
[0015]
[0016]
[0017]
[0018] The value range of the above t is [0, T - 1]; and then the adversarial sample obtained from the Tth perturbation is obtained.
[0019] Beneficial effects: The present invention directly transfers the perturbation space from the pixel space to the feature space to improve the efficiency of destroying important features, and then obtains the corresponding adversarial sample by means of feature inversion of the perturbed features. The present invention solves the problem of low efficiency of feature perturbation in traditional feature-level attacks, and makes the generated adversarial sample have stronger transferability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a schematic diagram of the overall framework of the present invention;
[0021] Figure 2 is the original image in the embodiment;
[0022] Figure 3 is the adversarial sample image generated by adopting the technical solution of the present invention in the embodiment;
[0023] Figure 4 is a schematic diagram of the overall framework of the prior art solution. DETAILED DESCRIPTION OF THE INVENTION
[0024] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.
[0025] Verification conclusion: For the classification model , the existing feature-level attacks include the following optimization problems.
[0026] The optimization problem of FIA is:
[0027] (1)
[0028] Among them, represents the dot product operation, represents the inner product operation, represents the aggregated gradient;
[0029] (2)
[0030] Among them, represents the number of ensembles, represents a binary matrix that satisfies the distribution, represents the image obtained by performing a dot product operation on the original image and the matrix, represents the logits output of the y-th dimension, represents the probability of performing random pixel dropout. Let , then the optimization problem of FIA can finally be rewritten as: (3)
[0031] The optimization problem of NAA is:
[0032] (4)
[0033] represents the value of the -th neuron in
[0034] (5)
[0035] Among them, represents the number of ensembles, represents the softmax output of the y-th dimension, , represents the all-black image, and are both linear functions that map to themselves, that is, , then formula (4) is rewritten as:
[0036] (6)
[0037] This is equivalent to: (7)
[0038] Let , the optimization problem of NAA is finally rewritten as: (8)
[0039] The optimization problem of BFA is: (9)
[0040] where, , represents the fitted image, is set to 1, making the optimization problem become: (10)
[0041] Since is a constant term independent of the optimization parameters, formula (10) is rewritten as:
[0042] (11)
[0043] Let , the optimization problem of BFA is finally rewritten as: (12)
[0044] In summary, the optimization problem of the existing feature-level attack is summarized as formula (12).
[0045] Here, formula (12) is transformed, and the transformation result is shown in formula (13):
[0046] (13)
[0047] where, represents the perturbation step size, which is a constant, so this transformation holds.
[0048] Then, formula (13) is transformed using the cosine theorem, and the result is shown in formula (14);
[0049] (14)
[0050] Because is a constant term, formula (14) is transformed into formula (15);
[0051] (15)
[0052] Specifically, maximizing means that the feature-level attack maximally destroys the intermediate-layer features of the surrogate model, while minimizing means that the feature-level attack tends to maximally destroy the intermediate-layer features of the surrogate model along the direction of .
[0053] The above verification process of the conclusion proves mathematically, by analyzing the loss functions of existing feature-level attacks, that existing feature-level attacks actually expect to perturb features along the direction of feature importance in the feature space.
[0054] That is to say, the effect of perturbing multiple times in the pixel space by existing feature-level attacks (as Figure 4 shown) is the same as the effect of perturbing features only once along the direction of feature importance in the feature space. This inefficient perturbation in the pixel space limits their improvement of adversarial transferability. To solve the problem of low efficiency of existing feature perturbations, the present invention proposes an attack method P2FA from the pixel space to the feature space. Specifically, according to feature importance, the perturbation space is directly transferred from the pixel space to the feature space to improve the efficiency of destroying important features, and then the perturbed features are used to obtain corresponding adversarial samples through the feature inversion method.
[0055] As Figure 1 shown, for the adversarial sample generation method of destroying intermediate layer features of the present invention, the classification model is set as , for this classification model, the adversarial sample generation method P2FA is used to perturb the intermediate layer features of the classification model along the direction with a step size of to form ; where and represent the input original image and the corresponding true label respectively; represents the feature map of the th intermediate layer; when performing the above perturbation
[0056] T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into local optima. The specific process is as follows:
[0057] ;
[0058] ;
[0059] where is the decay factor in momentum, is the perturbation step size, is the feature importance; the iterative formula of is as follows: ; ;
[0060] Finally, the perturbed features are inverted to the pixel space using the feature inversion algorithm to obtain the corresponding adversarial sample , that is, summarize the feature inversion as the following optimization problem:
[0061] ;
[0062] Invert the perturbed features to the pixel space through the feature inversion algorithm to obtain adversarial samples .
[0063] The present invention analyzes the deficiencies existing in the algorithms of existing feature-level attacks, focuses on solving the problem of low efficiency of perturbing features in existing feature-level attacks, and finally generates more transferable adversarial samples.
[0064] Before perturbing the intermediate layer features of the classification model T times in this embodiment, first initialize the adversarial samples when not iterated and the intermediate layer feature update direction , and the formulas are as follows:
[0065] ; ; ;
[0066] Then, during the above T perturbations, the adversarial samples and intermediate layer feature update directions generated at the th iteration are respectively denoted as , ,
[0067]
[0068]
[0069]
[0070] The value range of the above t is [0, T - 1]; and then the adversarial samples obtained by the Tth perturbation are obtained .
[0071] The above adversarial sample generation method can be represented by the P2FA algorithm.
[0072]
[0073] To verify the technical effect of the present invention, this embodiment compares the attack success rates of the technical solution P2FA of the present invention and existing feature-level attacks (FIA, NAA, BFA).
[0074] Table 1 Comparison of the attack success rates with existing feature-level attacks (FIA, NAA, BFA).
[0075]
[0076] Table 1 uses Inception-v3 (Inc-v3), Inception-v4 (Inc-v4), Inception-ResNet-v2 (IncRes-v2), and ResNet-152 (Res-152) as source models to generate adversarial samples to attack different target models. Figure 2 is the original clean image input in this embodiment, Figure 3 is the corresponding adversarial sample generated using the P2FA technical solution of the present invention. It can be seen that the generated adversarial sample is almost visually indistinguishable from the original clean image. However, the generated adversarial sample can mislead the classification model.
Claims
1. A method for generating adversarial samples that destroys intermediate layer features, characterized in that: Set the classification model to For this classification model, the adversarial sample generation method P2FA is used to generate features along the feature importance in the feature space. Direction in steps Perturb the intermediate layer features of the classification model to obtain the perturbed intermediate layer features ; in, and Represent the input original image and the corresponding true label respectively; Indicates Feature maps of the middle layer; When performing the above perturbations T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into the local optimum. The specific process is as follows: ; ; in, is the decay factor in momentum, is the perturbation step length, For the The update direction of the intermediate layer features at the iteration, || ||2 represents the L2 norm; is the feature importance, The iteration formula is as follows: ; ; n is the number of aggregations to obtain feature importance, ranging from 1 to N. represents the fitting image at the nth iteration, represents the disturbance size, represents the cross entropy loss; Finally, the perturbed features Invert to pixel space to obtain the corresponding adversarial sample , that is, the feature inversion is summarized as the following optimization problem: ; in, Represents the square of the L2 norm.
2. The method for generating adversarial samples that destroy intermediate layer features according to claim 1, characterized in that: Before perturbing the intermediate layer features of the classification model T times, initialize the adversarial sample before iteration And the intermediate layer feature update direction , the formula is as follows: ; ; ; Then in the above T disturbance process, The adversarial samples generated at the iteration and the update direction of the intermediate layer features are recorded as , , ; ; ; The value range of t is [0, T-1]; then we get the adversarial sample obtained by the T-th perturbation. .
Citation Information
Patent Citations
Multi-modal sentiment classification method based on isomorphism and heterogeneity dynamic information interaction
CN116010595A
High-mobility adversarial sample generation method and system
CN116011558A
Generation method of mobility confrontation sample and black box attack method
CN118052273A